Papers with model assessment

3 papers
Are Large Language Models Economically Viable for Industry Deployment? (2026.acl-industry)

Copied to clipboard

Challenge: Generative AI is increasingly deployed in healthcare, financial analytics, and conversational automation.
Approach: They propose a framework that evaluates large language models across their full lifecycle on legacy GPUs.
Outcome: The proposed framework evaluates LLMs across their full lifecycle on legacy GPUs.
Dynabench: Rethinking Benchmarking in NLP (2021.naacl-main)

Copied to clipboard

Challenge: Dynabench is an open-source platform for dynamic dataset creation and model benchmarking.
Approach: They propose an open-source platform for dynamic dataset creation and model benchmarking.
Outcome: The proposed platform can be used to create models that fail on simple challenges and falter in real-world scenarios.
SpeechAlign: A Framework for Speech Translation Alignment Evaluation (2024.lrec-main)

Copied to clipboard

Challenge: Speech-to-Speech and Speech- to-Text translation are currently dynamic areas of research.
Approach: They propose a framework to evaluate source-target alignment in speech models . they introduce a speech gold alignment dataset and introduce two new metrics .
Outcome: The proposed framework evaluates source-target alignment quality within speech models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations